Distribution by Document Size

نویسنده

  • Andrew Kane
چکیده

Search engines split large datasets across multiple machines using document distribution. Documents are typically distributed randomly to produce good load balancing. We propose that documents be distributed by their size instead. This can make load balancing more difficult, but it produces immediate improvements in both index size and query throughput. To support our proposal, we show improvements to an in-memory conjunctive list intersection system running on the GOV2 dataset using simple16 compression combined with either skips or bitvectors. While our list intersection system does not implement ranking, we also expect significant performance improvements from using document size distribution in ranking based search systems. In addition, implementations can be adapted to produce further performance improvements that exploit the document size distribution. We present some examples that apply to a full ranking based search engine and leave their verification for future work.

برای دانلود متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

منابع مشابه

یک مدل موضوعی احتمالاتی مبتنی بر روابط محلّی واژگان در پنجره‌های هم‌پوشان

A probabilistic topic model assumes that documents are generated through a process involving topics and then tries to reverse this process, given the documents and extract topics. A topic is usually assumed to be a distribution over words. LDA is one of the first and most popular topic models introduced so far. In the document generation process assumed by LDA, each document is a distribution o...

متن کامل

Document Analysis And Classification Based On Passing Window

In this paper we present Document analysis and classification system to segment and classify contents of Arabic document images. This system includes preprocessing, document segmentation, feature extraction and document classification. A document image is enhanced in the preprocessing by removing noise, binarization, and detecting and correcting image skew. In document segmentation, an algorith...

متن کامل

رفع اعوجاج هندسی متون به‌کمک اطلاعات هندسی خطوط متن

Document images produced by scanners or digital cameras usually have photometric and geometric distortions. If either of these effects distorts document, recognition of words from such a document image using OCR is subject to errors. In this paper we propose a novel approach to significantly remove geometric distortion from document images. In this method first we extract document lines from do...

متن کامل

An Analysis of Ministry of Education’s Strategic Plans Based on Favorable Components of English Language Teaching Using Shannon’s Entropy

The present research aims to analyze the content of Ministry of Education’s strategic plans (the Fundamental Reform Document of Education, the Comprehensive National Scientific Plan and the National Curriculum Document) based on Shannon's entropy regarding the favorable components of teaching English. The contents of the Fundamental Reform Document of Education, the Comprehensive National Scien...

متن کامل

Quantum dots of CdS synthesized by micro-emulsion under ultrasound: size distribution and growth kinetics

Quantum dots of CdS with hexagonal phase were prepared at relatively low temperature (60 oC) and short time by micro-emulsion (O/W) under ultrasound. This study was focused on the particle size distribution and the growth kinetics. The particle size distribution obtained from the optical absorption edge. It was relatively symmetrical with sonication time. In addition, an agreement was observed ...

متن کامل

ذخیره در منابع من


  با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید

برای دانلود متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

عنوان ژورنال:

دوره   شماره 

صفحات  -

تاریخ انتشار 2014